Papers with statistical analyses
EfficientOCR: An Extensible, Open-Source Package for Efficiently Digitizing World Knowledge (2023.emnlp-demo)
Copied to clipboard
| Challenge: | Existing OCR engines fail to provide accurate, cost-effective and sample-efficient character recognition for public domain documents. |
| Approach: | EffOCR is an open-source optical character recognition package that is accurate, cheap to deploy and sample efficient to customize to novel collections, languages, and character sets. |
| Outcome: | EffOCR model trains character retrieval problem and scales to novel collections, languages, and character sets. |
Does Generative AI speak Nigerian-Pidgin?: Issues about Representativeness and Bias for Multilingualism in LLMs (2025.findings-naacl)
Copied to clipboard
| Challenge: | Nigeria is a multilingual country with 500+ languages. |
| Approach: | They propose to use a pidgin and a creole to analyze the pidgins of Nigeria . they also use machine translation to analyze their results . |
| Outcome: | The results show that the two pidgins do not represent each other and are hard to teach . the results show the pidgin varieties are underrepresented in Generative AI . |
Replicating and Extending “Because Their Treebanks Leak”: Graph Isomorphism, Covariants, and Parser Performance (2021.acl-short)
Copied to clipboard
| Challenge: | a small sample size and unreliable results suggest a correlation between parser performance and graph isomorphism is not observed in the wild. |
| Approach: | They propose to replicate a study which found graph isomorphism is a non-trivial variable . they also bin sentences by length and find correlation between parser performance and isopathism disappears . |
| Outcome: | The results show that the original analysis was unreliable and had methodological issues . the study also bin sentences by length and shows that the correlation between parser performance and graph isomorphism disappears when controlling for covariants. |
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented. |
| Approach: | They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods. |
| Outcome: | The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents. |
ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media (2025.emnlp-main)
Copied to clipboard
Aakash Kumar Agarwal, Saprativa Bhattacharjee, Mauli Rastogi, Jemima S. Jacob, Biplab Banerjee, Rashmi Gupta, Pushpak Bhattacharyya
| Challenge: | Almost 50% of depression patients face the risk of going into relapse. |
| Approach: | They propose to validate a social media dataset on depression relapse using cognitive theories of depression. |
| Outcome: | The first clinically validated social media dataset focused on depression relapse comprises 204 Reddit users annotated by mental health professionals. |